Papers with sentence fusion

14 papers
Deep Learning Approaches to Text Production (N18-6)

Copied to clipboard

Challenge: Text production is a key component of many NLP applications . Claire Gardent is based in France and is pursuing research in text production .
Approach: This tutorial will cover the fundamentals and state-of-the-art research on neural models for text production.
Outcome: This tutorial will cover the fundamentals and the state-of-the-art research on neural models for text production.
Analyzing Sentence Fusion in Abstractive Summarization (D19-54)

Copied to clipboard

Challenge: Abstractive summarization systems struggle to combine information from multiple sources, resulting in poor grammar and incorrect facts.
Approach: They analyze the outputs of five abstractive summarization systems and examine their grammatical accuracy and faithfulness.
Outcome: The proposed summarization systems are able to combine information from multiple sources, but they often fail to remain faithful to the original document.
Abstractive Unsupervised Multi-Document Summarization using Paraphrastic Sentence Fusion (C18-1)

Copied to clipboard

Challenge: a new method for abstractive summarization is being developed for document summarizing . abstractive methods require extensive natural language generation to rewrite the sentences .
Approach: They propose an unsupervised abstractive summarization system in multi-document context . they use a paraphrastic sentence fusion model which performs sentence synthesis and paraphrazing .
Outcome: The proposed model improves information coverage and abstractiveness of generated sentences.
ASDOT: Any-Shot Data-to-Text Generation with Pretrained Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to data-to-text generation require limited training examples . a data-based approach is based on a set of pre-trained language models with optional finetuning.
Approach: They propose a data-to-text generation task that makes use of any given (or no) examples.
Outcome: The proposed approach improves on baselines on a dataset with zero/few/full-shot settings.
Event Graph based Sentence Fusion (2021.emnlp-main)

Copied to clipboard

Challenge: Sentence fusion is a conditional generation task that merges related sentences into a coherent text.
Approach: They propose to build an event graph from the input sentences to capture related events in a structured way and use the constructed event graph to guide sentence fusion.
Outcome: The proposed method achieves state-of-the-art on two datasets . it is based on the input sentences and shows that it is effective .
Learning to Fuse Sentences with Transformers for Summarization (2020.emnlp-main)

Copied to clipboard

Challenge: Abstractive summarization systems that fuse sentences are not rewarded for correctly fusing sentences.
Approach: They propose to leverage the knowledge of points of correspondence between sentences to enhance their ability to fuse sentences.
Outcome: The proposed algorithms improve the ability of the proposed summarization systems to fuse sentences and show that they can fuse sentences in a way that retains the original meaning.
DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion (N19-1)

Copied to clipboard

Challenge: Existing datasets for sentence fusion are small and insufficient for training modern neural models.
Approach: They propose a method for automatically-generating fusion examples from raw text . they apply their method to Wikipedia and Sports articles to generate fusion models .
Outcome: The proposed method improves performance on WebSplit when viewed as a sentence fusion task.
Seq2Edits: Sequence Transduction Using Span-level Edit Operations (2020.emnlp-main)

Copied to clipboard

Challenge: Seq2Edits is an open-vocabulary approach to sequence editing for natural language processing tasks with a high degree of overlap between input and output texts.
Approach: They propose an open-vocabulary approach to sequence editing for NLP tasks with a high degree of overlap between input and output texts.
Outcome: The proposed approach speeds up inference by up to 5.2x compared to full sequence models . it improves explainability by associating each edit operation with a human-readable tag.
Encode, Tag, Realize: High-Precision Text Editing (D19-1)

Copied to clipboard

Challenge: Neural sequence-to-sequence models provide a powerful framework for learning to translate source texts into target texts.
Approach: They propose a sequence tagging approach that casts text generation as a text editing task.
Outcome: The proposed model outperforms strong seq2seq models on sentence fusion, sentence splitting, abstractive summarization, and grammar correction tasks and achieves state-of-the-art performance.
Phrase-Level Localization of Inconsistency Errors in Summarization by Weak Supervision (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for evaluating inconsistency in summarization are limited . a recent study found that more than 30% of summarized summaries are inconsistent with the source documents .
Approach: They propose a method for localizing inconsistency errors in summarization using a synthetic dataset that contains factual errors likely to be produced by a common language processor.
Outcome: The proposed method detects factual errors more accurately than existing weakly supervised methods . the proposed model also detects errors in original sentences more accurately .
Dissecting Generation Modes for Abstractive Summarization Models via Ablation and Attribution (2021.acl-long)

Copied to clipboard

Challenge: Abstractive summarization models have made great strides in recent years, but little is known about how they actually form summaries and how to understand where their decisions come from.
Approach: They propose a two-step method to interpret summarization model decisions by categorizing each decoder decision into one of several generation modes.
Outcome: The proposed method can identify phrases the summarization model has memorized and determine where in the training pipeline this memorization happened, and study complex generation phenomena on a per-instance basis.
Improving Iterative Text Revision by Learning Where to Edit from Other Revision Tasks (2022.emnlp-main)

Copied to clipboard

Challenge: Iterative text revision improves text quality by fixing grammatical errors, rephrasing for better readability or contextual appropriateness.
Approach: They propose to build an end-to-end text revision system that can iteratively generate helpful edits by explicitly detecting editable spans with their corresponding edit intents.
Outcome: The proposed system outperforms baselines on other text revision tasks and human evaluations.
Unsupervised Text Style Transfer with Padded Masked Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for style transfer are difficult to obtain and require substantial amounts of parallel training examples to work well.
Approach: They propose an unsupervised method for style transfer that uses masked language models to find the text spans where the two models disagree the most in terms of likelihood.
Outcome: The proposed method performs competitively in a fully unsupervised setting and improves accuracy in low-resource settings by over 10 percentage points when pre-training on silver training data generated by Masker.
Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional Operations (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to learn sentence embeddings do not capture the semantic similarity of sentences.
Approach: They propose a framework that integrates compositional sentence operations into the embedding space and optimizes operator networks and a bottleneck encoder-decoder model to produce meaningful and interpretable sentence embeddables.
Outcome: The proposed framework improves the interpretability of sentence embeddings on four textual generation tasks while maintaining strong performance on traditional semantic similarity tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations